Papers with NLP research

65 papers
PhoNLP: A joint multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and dependency parsing (2021.naacl-demos)

Copied to clipboard

Challenge: PhoNLP is a multi-task learning model for joint Vietnamese part-of-speech (POS) tagging, named entity recognition (NER) and dependency parsing.
Approach: They propose a multi-task learning model for Vietnamese part-of-speech tagging, named entity recognition and dependency parsing that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently.
Outcome: The proposed model outperforms a single-task learning approach that fine-tunes the pre-trained Vietnamese language model PhoBERT for each task independently.
Security Challenges in Natural Language Processing Models (2023.emnlp-tutorial)

Copied to clipboard

Challenge: Large-scale natural language processing models are vulnerable to security issues, such as backdoor attacks, private data leakage, and imitation attacks.
Approach: They will dive into three emerging security issues in NLP research, i.e., backdoor attacks, private data leakage, and imitation attacks.
Outcome: This tutorial will cover three emerging security issues in NLP research, i.e., backdoor attacks, private data leakage, and imitation attacks.
Aggregating and Learning from Multiple Annotators (2021.eacl-tutorials)

Copied to clipboard

Challenge: NLP is based on high-quality annotated datasets, but many other tasks have unique characteristics not considered by standard models of annotation.
Approach: tutorial aims to connect NLP researchers with state-of-the-art aggregation models for canonical language annotation tasks.
Outcome: This tutorial aims to connect NLP researchers with state-of-the-art aggregation models for a diverse set of canonical language annotation tasks.
Writing Code for NLP Research (D18-3)

Copied to clipboard

Challenge: upcoming workshop on open source software for NLP aims to share best practices for writing code for Nl research . participants will learn how to write research code that facilitates good science and easy experimentation .
Approach: this tutorial aims to share best practices for writing code for NLP research . participants will learn how to write research code that facilitates good science and easy debugging .
Outcome: the workshop on open source software for NLP aims to share best practices for writing code for Nl research . participants will learn how to write research code that facilitates good science and easy experimentation .
LingConv: An Interactive Toolkit for Controlled Paraphrase Generation with Linguistic Attribute Control (2025.emnlp-demos)

Copied to clipboard

Challenge: LINGCONV is an interactive toolkit for controllable text generation . it allows fine-grained control over 40 specific linguistic attributes spanning lexical, syntactic, and discourse dimensions.
Approach: They propose a toolkit for paraphrase generation that allows finegrained control over 40 specific linguistic attributes.
Outcome: The toolkit is available at https://mohdelgaar-lingconv.hf.space, with a demo video at https:youtu.be/wRBJEJ6EALQ.
Deep Learning for Conversational AI (N18-6)

Copied to clipboard

Challenge: Spoken Dialogue Systems (SDS) have great commercial potential . the advent of deep learning has led to significant advances in this area of NLP research .
Approach: This tutorial will introduce researchers to the pipeline framework for modelling goal-oriented dialogue systems.
Outcome: This tutorial will familiarise researchers with the latest advances in spoken dialogue systems . the aim of the course is to encourage dialogue research in the NLP community .
A Two-Sided Discussion of Preregistration of NLP Research (2023.eacl-main)

Copied to clipboard

Challenge: et al. (2021) suggest NLP research should adopt preregistration to prevent fishing expeditions and promote publication of negative results.
Approach: et al. suggest NLP research should adopt preregistration to prevent fishing expeditions and promote publication of negative results.
Outcome: The proposed approach solves many methodological problems with NLP research.
Text Characterization Toolkit (TCT) (2022.aacl-demo)

Copied to clipboard

Challenge: Text Characterization Toolkit (TCT) is a tool that researchers can use to study characteristics of large datasets.
Approach: They propose a text characterization toolkit that researchers can use to study characteristics of large datasets.
Outcome: The proposed tool can be used to study characteristics of large datasets and to understand the influence of attributes on models’ behaviour.
GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture (2025.acl-demo)

Copied to clipboard

Challenge: generative large language models have become crucial for modern NLP research and applications across multiple languages.
Approach: They introduce the GigaChat family of Russian LLMs, available in various sizes . they evaluate their performance on Russian and English benchmarks and compare them with multilingual analogs .
Outcome: The proposed model family is available in various sizes and is tested on Russian and English benchmarks.
Should we find another model?: Improving Neural Machine Translation Performance with ONE-Piece Tokenization Method without Model Modification (2021.naacl-industry)

Copied to clipboard

Challenge: Recent studies using pretrain-finetuning approach have achieved state-of-the-art (SOTA) performance in many natural language processing tasks.
Approach: They propose a new tokenization method that combines morphology-considered subword tokenization and vocabulary methods to address this limitation.
Outcome: The proposed method can be used without modifying the model structure.
Deep Pivot-Based Modeling for Cross-language Cross-domain Transfer with Minimal Guidance (D18-1)

Copied to clipboard

Challenge: a framework for cross-domain and cross-language transfer has hardly been explored . cross-linguistic and cross language transfer methods are used for multilingual applications .
Approach: They propose a framework that builds on pivot-based learning, structure-aware Deep Neural Networks and bilingual word embeddings to train a model on labeled data from one language pair.
Outcome: The proposed model outperforms existing models even when trained in the lazy setup . the proposed model can be applied to nine English-German and nine English - french domain pairs without retraining .
You Only Need Attention to Traverse Trees (P19-1)

Copied to clipboard

Challenge: Recent research has focused on sentence representations.
Approach: They propose a tree-based model that captures phrase-level syntax and word-level dependencies by doing recursive traversal with attention.
Outcome: a new model captures phrase-level syntax and word-level dependencies with attention.
“John is 50 years old, can his son be 65?” Evaluating NLP Models’ Understanding of Feasibility (2023.eacl-main)

Copied to clipboard

Challenge: Recent work has found that large-scale language models lack commonsense reasoning ability . a dataset evaluating large-level language models is needed to evaluate their understanding of feasibility .
Approach: They propose a question-answering dataset that tests understanding of feasibility . they propose to use commonsense reasoning to reason about when an action is feasible .
Outcome: The proposed dataset shows that state-of-the-art models struggle to answer feasibility questions correctly.
Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research (2026.tacl-1)

Copied to clipboard

Challenge: Analogical reasoning is an essential aspect of human cognition, says aaron eliotta . eelisa e. sabet: some have argued that analogy is central to the human cognitive experience .
Approach: They summarize key theories about the processes underlying analogical reasoning from the cognitive science literature and relate it to current research in natural language processing.
Outcome: The proposed approaches are relevant for several major challenges in natural language processing, not directly related to analogy solving.
ExplainaBoard: An Explainable Leaderboard for NLP (2021.acl-demo)

Copied to clipboard

Challenge: Using leaderboards, researchers can track the performance of various systems on various NLP tasks.
Approach: They propose a new conceptualization and implementation of NLP evaluation using a leaderboard.
Outcome: The ExplainaBoard is an evaluation tool for natural language processing (NLP) it covers more than 400 systems, 50 datasets, 40 languages, and 12 tasks.
FALTE: A Toolkit for Fine-grained Annotation for Long Text Evaluation (2022.emnlp-demos)

Copied to clipboard

Challenge: Existing tools to evaluate long text outputs are lacking in the field of NLP . human rating and error analysis remains a crucial component for any evaluation of long text generation.
Approach: They propose a web-based toolkit to collect fine-grained error annotations for long texts . they use a taxonomy to identify errors and assign them to text spans .
Outcome: The proposed tool can be used to evaluate the coherence of long generated summaries.
Preregistering NLP research (2021.naacl-main)

Copied to clipboard

Challenge: Preregistration refers to specifying what you are going to do, and what you expect to find in your study, before carrying out the study.
Approach: They propose to use preregistration to specify what you are going to do and what you expect to find in your study, and propose several preregistrations questions for different kinds of studies.
Outcome: The proposed preregistration form could provide firmer grounds for slow science in NLP research.
Towards the First NLP Benchmark for Ladin - an Extremely Low-Resource Language (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) are limited in low-resource languages due to lack of labeled training data.
Approach: They propose to use Ladin as a model for sentiment analysis and question answering by incorporating Italian data into machine translation training.
Outcome: The proposed method improves on existing Italian–Ladin translation baselines.
Deep Temporal-Recurrent-Replicated-Softmax for Topical Trends over Time (N18-1)

Copied to clipboard

Challenge: a novel topic model is proposed to allow topical trends to be captured in temporal collections of documents.
Approach: They propose a novel unsupervised neural dynamic topic model where topics are influenced by topic discovery over time.
Outcome: The proposed model shows better generalization, topic interpretation, evolution and trends compared to state-of-the-art models .
XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing datasets are often informed by established research directions in the NLP community.
Approach: They propose a benchmark to evaluate the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
Outcome: The proposed benchmark evaluates the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies including ASR, OCR, MT, and information access tasks.
The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing (P18-1)

Copied to clipboard

Challenge: Statistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental.
Approach: They propose a protocol for statistical significance test selection in NLP setups . they propose he proposes a survey of the most relevant tests to help guide the protocol .
Outcome: The proposed protocol includes a survey of the most relevant tests.
MLD-EA: Check and Complete Narrative Coherence by Introducing Emotions and Actions (2025.coling-main)

Copied to clipboard

Challenge: Existing studies focus on summarization and question-answering tasks, but neglect logical coherence within stories.
Approach: They propose a model that leverages large language models to identify narrative gaps and generate coherent sentences that integrate seamlessly with the story’s emotional and logical flow.
Outcome: The proposed model enhances narrative understanding and story generation, highlighting LLMs’ potential as effective logic checkers in story writing with logical coherence and emotional consistency.
Argument Quality Assessment in the Age of Instruction-Following Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Argument quality assessment is critical for opinion formation, decision making, writing education, and the like.
Approach: They propose to use large language models to leverage knowledge across contexts to enable a much more reliable assessment.
Outcome: The proposed approach improves the quality of argumentation and the ability to leverage knowledge across contexts.
SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages (2024.findings-acl)

Copied to clipboard

Challenge: SemRel datasets are annotated by native speakers across 13 languages . they are used to characterise the relationship between two units of text .
Approach: They propose to use a semantic relatedness dataset to measure the degree of semantic textual relatedness between sentences in Afrikaans, Algerian Arabic, Amharic, English, Hausa, Hindi, Indonesian, Kinyarwanda, Marathi, Moroccan Arabic, Modern Standard Arabic, Spanish, and Telugu.
Outcome: The proposed datasets are annotated by native speakers across 13 languages and represent the semantic relatedness of 13 languages.
A Survey of Race, Racism, and Anti-Racism in NLP (2021.acl-long)

Copied to clipboard

Challenge: despite inextricable ties between race and language, little work has considered race in NLP research and development.
Approach: They survey 79 papers from the ACL anthology that mention race . they find race has been siloed as a niche topic and ignored in many NLP tasks . authors call for inclusion and racial justice in NLP research practices .
Outcome: The findings highlight the need for inclusion and racial justice in NLP research practices.
Collaborative Performance Prediction for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are one of the most important AI research powered by largescale parameters, high computational resources, and massive training data.
Approach: They propose a framework that leverages historical performance of large language models and other design factors to improve prediction accuracy.
Outcome: The proposed framework surpasses scaling laws in predicting performance of large language models . it also facilitates a detailed analysis of factor importance, an area previously overlooked .
Towards Climate Awareness in NLP Research (2022.emnlp-main)

Copied to clipboard

Challenge: Increasing focus on efficient AI and NLP research lacks systematic climate reporting guidelines . a proposed model card would be practical with limited information about experiments and the underlying computer hardware.
Approach: They propose a model card that is practically usable with limited information about experiments and the underlying computer hardware.
Outcome: The proposed model card would be usable with limited information about experiments and the underlying computer hardware.
COVID-19 Named Entity Recognition for Vietnamese (2021.naacl-main)

Copied to clipboard

Challenge: a new dataset is being developed to help fight the COVID-19 pandemic . the dataset is annotated for the named entity recognition task with newly-defined entity types .
Approach: They present the first manually-annotated COVID-19 domain-specific dataset for Vietnamese . their dataset is annotated for the named entity recognition task with newly-defined entity types .
Outcome: The proposed dataset is the first manually-annotated COVID-19 domain-specific dataset for Vietnamese.
Is NLP Ready for Standardization? (2022.findings-emnlp)

Copied to clipboard

Challenge: a number of scientific fields, including telecommunications, networks and multimedia, lack standards in the field of NLP.
Approach: They propose to examine how NLP lacks standards and how that can impact society, industry and regulations.
Outcome: The proposed standards examine the needs of NLP researchers and industry . they argue that the lack of standards can impact the field, society and industry.
ZmBART: An Unsupervised Cross-lingual Transfer Framework for Language Generation (2021.findings-acl)

Copied to clipboard

Challenge: Recent advances in NLP focus on large annotated training data.
Approach: They propose an unsupervised framework that does not use parallel or pseudo-parallel/back-translated data.
Outcome: The proposed framework does not use parallel or pseudo-parallel/back-translated data.
Academic-Industrial Perspective on the Development and Deployment of a Moderation System for a Newspaper Website (L18-1)

Copied to clipboard

Challenge: a system that supports the moderation of user comments on a large newspaper website is described in this paper.
Approach: They describe an approach and experiences from the development, deployment and usability testing of a natural language processing and information retrieval system that supports the moderation of user comments on a large newspaper website.
Outcome: The proposed system supports the moderation of user comments on a large newspaper website.
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Cantonese has scant representation in NLP research, especially compared to other languages from similarly developed regions.
Approach: They propose to evaluate Cantonese LLM performance in factual generation, mathematical logic, complex reasoning, and general knowledge in Cantonesian.
Outcome: The proposed models will evaluate Cantonese's performance in factual generation, mathematical logic, complex reasoning, and general knowledge in Cantone.
LEGAL-BERT: The Muppets straight out of Law School (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing guidelines for pre-training and fine-tuning do not always generalize well in the legal domain.
Approach: They propose to use BERT out of the box, adapt it by additional pre-training on domain-specific corpora, and pre-train it from scratch on domains.
Outcome: The proposed strategies are: use the original BERT out of the box, adapt it by additional pre-training on domain-specific corpora, and pre-train it from scratch on domain specific corpors.
Give Me Convenience and Give Her Death: Who Should Decide What Uses of NLP are Appropriate, and on What Basis? (2020.acl-main)

Copied to clipboard

Challenge: a paper on automatic sentencing was a source of debate at EMNLP 2019 . paper examines whether particular datasets and tasks should be off-limits for NLP research .
Approach: They propose a neural model which performs structured prediction of individual charges laid against an individual and the prison term associated with each.
Outcome: The proposed model can predict the prison term associated with a given case on a large-scale dataset of real-world Chinese court cases.
How Good Is NLP? A Sober Look at NLP Tasks through the Lens of Social Impact (2021.findings-acl)

Copied to clipboard

Challenge: Recent years have seen many breakthroughs in natural language processing (NLP), transitioning it from a mostly theoretical field to one with many real-world applications.
Approach: They propose a moral philosophy definition of social good and a framework to evaluate the direct and indirect real-world impact of NLP tasks.
Outcome: The proposed framework evaluates the direct and indirect real-world impact of NLP tasks and adopts the methodology of global priorities research to identify priority causes for NLP research.
Beyond Fair Pay: Ethical Implications of NLP Crowdsourcing (2021.naacl-main)

Copied to clipboard

Challenge: Ethical considerations regarding the use of crowdworkers are limited to labor conditions . the Final Rule did not anticipate the use online crowdsourcing platforms for data collection .
Approach: They propose to reopen discussion regarding ethical use of crowdworkers in NLP research . they propose to use online crowdsourcing platforms to evaluate risk of harm .
Outcome: The proposed study identifies common scenarios where crowdworkers performing NLP tasks are at risk of harm.
A Neural Pairwise Ranking Model for Readability Assessment (2022.findings-acl)

Copied to clipboard

Challenge: Automatic Readability Assessment (ARA) is traditionally treated as a classification problem in NLP research.
Approach: They propose a neural ranking approach to automatic readability assessment (ARA) they propose 'neural' ranking methods that can be used to rank texts by reading level .
Outcome: The proposed approach performs well in monolingual single/cross corpus testing scenarios and achieves a zero-shot cross-lingual ranking accuracy of over 80% for both French and Spanish when trained on English data.
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies.
Approach: They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances .
Outcome: The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations.
Energy and Policy Considerations for Deep Learning in NLP (P19-1)

Copied to clipboard

Challenge: Recent advances in hardware and methodology for training neural networks have enabled significant accuracy improvements across many NLP tasks.
Approach: They quantify the approximate financial and environmental costs of training neural network models . they propose actionable recommendations to reduce costs and improve equity in NLP research .
Outcome: The proposed recommendations address the cost and environmental costs of training neural networks for NLP.
Systematic Inequalities in Language Technology Performance across the World’s Languages (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages.
Approach: They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Outcome: The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Data-Informed Global Sparseness in Attention Mechanisms for Deep Neural Networks (2024.lrec-main)

Copied to clipboard

Challenge: Attention pruning techniques have been developed to identify and exploit sparseness . previous work has taken pioneering steps to discover and explain the sparsity in attention patterns .
Approach: They propose a framework that observes attention patterns in a fixed dataset and generates a global sparseness mask.
Outcome: The proposed approach saves 90% of computations and maintains quality of results.
Phrase-level Self-Attention Networks for Universal Sentence Encoding (D18-1)

Copied to clipboard

Challenge: Phrase-level self-attention networks (PSAN) can capture context dependencies at the phrase level instead of the sentence level.
Approach: They propose to perform self-attention across words inside a phrase to capture context dependencies at the phrase level and use the gated memory updating mechanism to refine each word’s representation hierarchically with longer-term context dependency captured in a larger phrase.
Outcome: The proposed model can achieve state-of-the-art performance across a plethora of NLP tasks including binary and multi-class classification, natural language inference and sentence similarity.
COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline Generation (2025.acl-long)

Copied to clipboard

Challenge: COSMMIC is a multimodal, multilingual dataset featuring nine major Indian languages.
Approach: They propose a multimodal, multilingual multimodal multimodal dataset that integrates text, images and user feedback to enhance summarization.
Outcome: The proposed dataset is based on 4,959 article-image pairs and 24,484 reader comments with ground-truth summaries available in all included languages.
Towards large language model-based personal agents in the enterprise: Current trends and open problems (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing large language models (LLMs) are brittle to input changes and can produce inconsistent results for the same inputs.
Approach: They propose to use large language models to reason about complex goals and orchestrate a set of pluggable tools or APIs to accomplish a goal.
Outcome: The proposed use cases have many open problems in an exciting area of NLP research, such as trust and explainability, consistency and reproducibility, and the need for new metrics and benchmarks.
HellaSwag: Can a Machine Really Finish Your Sentence? (P19-1)

Copied to clipboard

Challenge: Existing commonsense models struggle to perform inferences that are trivial for humans, but are often misclassified by state-of-the-art models.
Approach: They propose a dataset that is adversarial to state-of-the-art commonsense reasoning and use it to build a model that is surprisingly robust.
Outcome: The proposed dataset is compared with existing models and scaled up towards a critical 'Goldilocks zone' wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models.
Co-Stack Residual Affinity Networks with Multi-level Attention Refinement for Matching Text Sequences (D18-1)

Copied to clipboard

Challenge: a long standing problem in NLP research is learning a matching function between two text sequences . a deep architecture for this task is proposed by a team of researchers .
Approach: They propose a new deep matching model using stacked recurrent encoders to learn affinity weights . they conduct extensive experiments on six well-studied text sequence matching datasets a plethora of applications are possible .
Outcome: The proposed model improves performance on six well-studied text sequence matching datasets.
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia (2022.acl-long)

Copied to clipboard

Challenge: There are more than 700 languages spoken in Indonesia, equal to 10% of the world's languages, second only to Papua New Guinea.
Approach: They focus on the languages spoken in Indonesia, the world's second most linguistically diverse nation, and the fourth most populous nation of the world.
Outcome: The proposed model is based on the languages spoken in Indonesia, the world's second-most linguistically diverse nation, with 273 million people spread over 17,508 islands.
Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess Hypotheses (2020.acl-main)

Copied to clipboard

Challenge: Empirical research in natural language processing has adopted a narrow set of principles for assessing hypotheses . alternative approaches to assess hypothese rely on p-value computation, which suffers from several known issues.
Approach: They propose to compare different methods for assessing hypotheses . they argue that practitioners should first decide their target hypothesis before choosing a method .
Outcome: The proposed method differs from other methods, but is not widely used in NLP . the proposed method is based on a p-value computation, but has a small gap in accuracy .
Regulation and NLP (RegNLP): Taming Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: polarization in AI safety and ethics debates are swaying political agendas on AI regulation and governance . regulation studies are rich source of knowledge on how to systematically deal with risk and uncertainty .
Approach: They argue that NLP research can benefit from proximity to regulatory studies . they argue that regulation studies should focus on linking scientific knowledge to regulatory processes .
Outcome: The proposed research space should focus on linking scientific knowledge to regulatory processes based on systematic methodologies.
CodE Alltag 2.0 — A Pseudonymized German-Language Email Corpus (2020.lrec-1)

Copied to clipboard

Challenge: unauthorized use of social media content as a data resource is often neglected . data privacy concerns are often overlooked in NLP research .
Approach: They propose an algorithm for the protection of personal data via pseudonymization by automatically recognizing privacy-sensitive stretches of text in UGC.
Outcome: The proposed algorithm protects personal data via pseudonymization on two hitherto non-anonymized German-language email corpora.
GeoSQA: A Benchmark for Scenario-based Question Answering in the Geography Domain at High School Level (D19-1)

Copied to clipboard

Challenge: SQA is an emerging application of NLP in the medical, geography, and legal domains.
Approach: They propose a dataset of 1,981 scenarios and 4,110 multiple-choice questions in geography domain at high school level.
Outcome: The proposed dataset consists of 1,981 scenarios and 4,110 multiple-choice questions in the geography domain at high school level.
A Major Obstacle for NLP Research: Let’s Talk about Time Allocation! (2022.emnlp-main)

Copied to clipboard

Challenge: Subpar time allocation has been a major obstacle for natural language processing research in recent years, argues a new position paper .
Approach: They propose to identify the biggest traps the NLP community falls into and suggest solutions to solve them.
Outcome: The authors outline multiple concrete problems together with their negative consequences and suggest remedies to improve the status quo.
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are useful for low-resource scenarios and time-restricted applications.
Approach: They propose a large-scale evaluation tool for large language models that uses prompts . they evaluate 720 prompt templates for open-source LLM-based metrics on MT and summarization datasets a 6.6M evaluations.
Outcome: The proposed model evaluates 720 prompt templates on machine translation and summarization datasets.
Perturbation Augmentation for Fairer NLP (2022.emnlp-main)

Copied to clipboard

Challenge: Unwanted and often harmful social biases are becoming more salient in NLP research.
Approach: They propose to train a neural perturbation model that rewrites demographic references in text to make them more fair.
Outcome: The proposed model outperforms heuristic alternatives on a large dataset of human annotated text perturbations.
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models (2023.acl-long)

Copied to clipboard

Challenge: Hundreds of studies have highlighted ethical issues in NLP models .
Approach: They propose to measure media biases in LMs trained on diverse data sources . they focus on hate speech and misinformation detection .
Outcome: The proposed methods quantify the fairness of downstream NLP models trained on politically biased LMs.
Curating Datasets for Better Performance with Example Training Dynamics (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to improve data quality but rely on data quantity to improve performance are not effective.
Approach: They propose a method for weighing the relative importance of examples in a dataset based on their Example Training dynamics (ETD) they propose an active learning approach for computing ETD during training rather than as a preprocessing step.
Outcome: The proposed method can be used to improve performance in in-distribution and out-of-distortion testing.
Verifying Annotation Agreement without Multiple Experts: A Case Study with Gujarati SNACS (2023.findings-acl)

Copied to clipboard

Challenge: a small fraction of the about 7,000 languages of the world have datasets or linguistic tools . linguistic datasets are a foundation of NLP research, but they are not always reliable . authors propose weak verifiers to help estimate dataset quality .
Approach: They propose four weak verifiers to help estimate dataset quality . they propose to use Gujarati as a low-resource language to test for dataset quality.
Outcome: The proposed methods concur with a double-annotation study in Gujarati.
Survey on Thai NLP Language Resources and Tools (2022.lrec-1)

Copied to clipboard

Challenge: Thai language is one of the under-resourced languages in the NLP domain, although it is spoken by approximately 70 million people globally.
Approach: They propose to use Thai language as an example to understand how NLP works and how it can be applied to Thai language.
Outcome: The results show that Thai NLP research has progressed over the past three decades, especially on upstream tasks such as tokenisation, but research on downstream tasks such syntactic parsing and semantic analysis is still limited.
Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP (2022.emnlp-main)

Copied to clipboard

Challenge: Textual adversarial samples are often misrepresented in research on security, evaluation, explainability, and data augmentation.
Approach: They propose to use adversarial samples to evaluate their methods on security tasks to demonstrate the real-world concerns rather than developing impractical methods.
Outcome: The proposed method has higher practical value than the current benchmark.
Interpreting Themes from Educational Stories (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in machine reading comprehension (MRC) have centered on literal comprehension, referring to the surface-level understanding of content.
Approach: They propose a dataset specifically designed for interpretive comprehension of educational narratives, providing corresponding well-edited theme texts.
Outcome: The proposed dataset spans genres and cultural origins and includes human-annotated theme keywords with varying levels of granularity.
IndoNLI: A Natural Language Inference Dataset for Indonesian (2021.emnlp-main)

Copied to clipboard

Challenge: XLM-R model outperforms other pre-trained models in annotated data.
Approach: They adapt the data collection protocol for MNLI and collect 18K sentence pairs annotated by crowd workers and experts.
Outcome: The proposed dataset outperforms other pre-trained models on the expert-annotated data.
A Comparison of Language Modeling and Translation as Multilingual Pretraining Objectives (2024.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models (PLMs) display impressive performances and have captured the attention of the NLP community.
Approach: They propose to compare multilingual pretraining objectives in a controlled methodological environment with multilingual models.
Outcome: The proposed model outperforms existing models in 6 languages and demonstrates that multilingual translation is an effective pretraining objective under the right conditions.
NLP Needs Diversity outside of ‘Diversity’ (2025.findings-emnlp)

Copied to clipboard

Challenge: a new position paper argues that diversity in NLP is concentrated on a small number of areas surrounding fairness .
Approach: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Outcome: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal Metaphors (2025.acl-long)

Copied to clipboard

Challenge: Metaphors are pervasive in communication, making them crucial for natural language processing.
Approach: They propose a multicultural multimodal metaphor dataset designed for cross-cultural studies of metaphor in Chinese and English.
Outcome: The proposed model improves metaphor comprehension across cultural backgrounds and cultural domains.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations